Back

DIGITAL HEALTH

SAGE Publications

Preprints posted in the last 7 days, ranked by how well they match DIGITAL HEALTH's content profile, based on 17 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Global Adoption of openEHR Clinical Data Repositories: A Vendor and Community Survey

Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.

2026-08-31 health informatics 10.64898/2026.08.27.26361529 medRxiv
Top 0.1%
7.7%
Show abstract

The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.

2
Preferences for receiving study results among pregnant women participating in a phase III clinical trial in Papua New Guinea.

Mengi, A.; Bagita-Vangana, M.; Tesine, P.; Laman, M.; Bolnga, J. W.; Ome-Kaius, M.; Kulimbao, J.; Mase, J.; Mal, L. S.; Mnjala, H.; Lee, G.; Cassidy-Seyoum, S. A.; Thriemer, K.; Unger, H. W.

2026-08-31 medical ethics 10.64898/2026.08.27.26361571 medRxiv
Top 0.2%
3.3%
Show abstract

Disseminating study results to participants is an ethical responsibility for researchers but remains uncommon in low- and middle-income countries, and participants preferences for receiving study results are poorly understood. This study examined study result dissemination preferences among pregnant women in a phase III malaria prevention trial in Papua New Guinea (PNG). Participants completed an interviewer-administered questionnaire (survey) assessing their interest in and motivation for receiving trial results and preferences for dissemination methods and content. Associations between participants characteristics and dissemination preferences were explored using multivariable logistic regression analysis. Of 1172 trial participants, 96.0% (1125/1172) completed the survey, and of these 99.6% (1121/1125) wanted to learn about the trial results. The main motivation factors driving participants interest were an acknowledgment of their contribution to research (51.7%; n=579) and a better understanding of the study (45.0%; n=505). Most participants (78.9%; n=884) wanted to learn about the trial findings through written summary and a group meeting with other participants at the nearest clinic (31.1%, n=349). Multivariable regression analysis indicated that participants from rural/peri-urban clinics were more likely to choose non-electronic media dissemination approaches such as a group meeting as compared to urban-dwelling participants. Frequently selected items (>50% of participants) for content included information regarding good results of the study, purpose of the study, medical treatment advances, results specific to me, and how study was conducted. There was heterogenicity in the preference for dissemination content: compared to urban clinics rural clinics are less likely to want to learn about how and why study was conducted and medical and scientific advances. Overall, the majority wanted to learn about trial results, highlighting the importance of integrating dissemination into research activities in PNG. Variation in preferences for mode and content of dissemination between study clinics suggests that dissemination activities could be tailored to local context and preferences.

3
Spectral and melanopic dose calibration of consumer see-through extended-reality glasses for controlled retinal photostimulation

Gaidica, M.; Rosengart, M.

2026-08-31 ophthalmology 10.64898/2026.08.26.26361398 medRxiv
Top 0.2%
3.3%
Show abstract

Light reaching the retina is a primary regulator of human circadian physiology, acting largely through melanopsin-expressing retinal ganglion cells with peak short-wavelength sensitivity. Delivering known, repeatable retinal doses outside the laboratory is difficult because conventional light sources leave viewing geometry, gaze, and ambient conditions uncontrolled. Consumer extended-reality (XR) glasses fix a bright binocular display in constant geometry relative to the eye, but their suitability as calibrated photic stimulators has not been established. Here we validate a commercial micro-OLED XR display (VITURE Luma Ultra) for controlled retinal photostimulation. A purpose-built host application renders exact 8-bit RGB stimuli while independently controlling hardware brightness and logging all intensity-determining state; spectral radiance was measured at the retinal position of a 3D-printed phantom head with an open-source miniature spectroradiometer, anchored to absolute units by a luminance transfer calibration. The blue primary peaks at 461 nm (FWHM 43 nm), is spectrally invariant across a >10-fold intensity range, and at maximum output delivers an estimated 299 lx melanopic equivalent daylight illuminance, above consensus daytime recommendations, while remaining roughly two orders of magnitude below photobiological safety limits. The red primary is visually effective with minimal melanopic drive (melanopic DER 0.10), enabling spectrally shifted evening stimulation. Unlike the immersive virtual-reality headsets previously used for calibrated light delivery, the see-through form factor preserves the wearer's view of the surroundings--relevant for clinical monitoring in supervised settings such as the intensive care unit. These results show that consumer XR glasses can serve as a dose-calibrated platform for wearable photostimulation using an open-source measurement chain, and provide groundwork for application-layer dose-response studies.

4
A Pragmatic Randomized Trial of an EHR-Integrated Generative AI Chart Summarization Tool for Ambulatory Clinicians

Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.

2026-08-31 health informatics 10.64898/2026.08.26.26361496 medRxiv
Top 0.2%
3.2%
Show abstract

BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.

5
When medical credentials conflict with stated accuracy: A factorial study of source credibility and answer revision in medical LLM interactions

Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.

2026-09-01 health informatics 10.64898/2026.08.28.26361634 medRxiv
Top 0.2%
2.5%
Show abstract

Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.

6
Are Frontier Large Language Models Safer Than Government-Backed Symptom Checkers for Clinical Self-Triage? A Standardised Vignette Evaluation

Chowdhury, A. R.; Chowdhury, B.

2026-09-02 health informatics 10.64898/2026.09.01.26361908 medRxiv
Top 0.2%
2.4%
Show abstract

Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.

7
Evaluating GPT-4o Model Proficiency and Clinical Reasoning for Antimicrobial Stewardship in Dentistry

Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.

2026-09-03 dentistry and oral medicine 10.64898/2026.09.01.26361980 medRxiv
Top 0.3%
1.8%
Show abstract

Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.

8
Understanding urgent blood-donor mobilisability: a cross-sectional online survey of digitally reachable adults in Ghana

Shen, H.; Agorinya, I. A.; Ayanore, M. A.; Brede, M.; Chapman, A.; Head, M.

2026-08-31 public and global health 10.64898/2026.08.27.26361538 medRxiv
Top 0.4%
1.7%
Show abstract

Introduction Safe and timely blood availability remains a major global health challenge, especially in low- and middle-income countries. Digital tools may accelerate donor contact, but digital reachability alone does not ensure that people will notice, trust and act on urgent requests to support blood donation efforts. We examined factors associated with anticipated engagement in digitally coordinated urgent blood-donor mobilisation among digitally reachable adults in Ghana. Methods We conducted a cross-sectional online survey from September 2025 to January 2026 across Ghana's 16 regions. Participants were recruited via Facebook advertising and snowball sampling. Factors associated with urgent blood-donor mobilisability were assessed under four criteria: high future-donation willingness; high willingness to install a trusted donation app; high willingness to respond to a trusted urgent-request; and high practical flexibility to leave current activities. Descriptive analyses and multivariable logistic regression examined prevalence and associated factors. Results Among 1,067 participants, 577 (54.1%) met all four criteria. Future-donation willingness (91.8%), trusted-app installation willingness (83.2%) and trusted-request response willingness (82.7%) were common, whereas practical flexibility was lower (66.6%). In the adjusted model, high formal health-system trust (adjusted OR (AOR) 3.95, 95% CI 2.08-7.50), high digital-response readiness (AOR 2.26, 1.66-3.08), previous donation (AOR 1.47, 1.08-2.01), high donation knowledge (AOR 1.42, 1.03-1.97) and willingness to donate to strangers were positively associated with high mobilisability. Women (AOR 0.60, 0.43-0.83), participants reporting a work-schedule barrier (AOR 0.43, 0.29-0.66) and those travelling over 30 min to the nearest healthcare facility at night (AOR 0.66, 0.45-0.96) had lower adjusted odds. Conclusions Digital reachability and stated donation willingness may overestimate the population pool available for emergency donation. Digital blood-donor solutions should consider verifiable health-system requests, account for response readiness and current availability, and connect willing individuals with accessible collection options and transport support where needed.

9
What Matters Most: A Multi-Stakeholder Study of Outcome Domains in Lower-Limb Prosthesis Use

Ahmed, M. E.; Karlsson-Brown, S.; Koufaki, P.; Ahmadi, M.; Mico-Amigo, E. M.

2026-09-03 rehabilitation medicine and physical therapy 10.64898/2026.08.31.26361544 medRxiv
Top 0.4%
1.7%
Show abstract

Purpose: Lower-limb prosthesis use involves interacting physical, psychosocial, and device-related outcomes that may not be fully captured by conventional clinical assessment. This study aimed to develop and evaluate a stakeholder-informed framework of outcome domains relevant to meaningful everyday prosthesis use. Materials and Methods: A mixed-methods participatory design comprised a structured synthesis of selected clinically relevant content from five established patient-reported outcome measures; semi-structured interviews and importance and actionability ratings with 18 contributors (12 prosthesis users, four clinicians, and two industrial partners); and integration of the synthesis, qualitative, and rating findings. Interview records were analysed using reflexive thematic analysis, and ratings were analysed descriptively. Results: The resulting framework comprised four interrelated domains: Mobility, Physical Function, Psychosocial Wellbeing, and Prosthesis Experience. Mobility showed the clearest convergence across stakeholder perspectives. Prosthesis users showed the largest importance actionability gap for Prosthesis Experience (4.5 vs 3.0), whereas clinicians showed the largest gap for Psychosocial Wellbeing (5.0 vs 3.0). Interviews highlighted day-to-day variability in prosthesis use and the influence of confidence, fatigue, comfort, environmental conditions, social context, and device usability. Conclusions: Meaningful outcome assessment in prosthetic rehabilitation should extend beyond mobility alone to consider physical function, psychosocial wellbeing, and prosthesis experience within everyday contexts. The proposed framework provides a stakeholder-informed foundation for multidimensional outcome assessment in prosthetic rehabilitation.

10
Default-filled outcome labels in a deployed cognitive-screening programme: an operator-level audit and the construction of twenty-four language-model arms

Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.

2026-09-02 health informatics 10.64898/2026.08.28.26361585 medRxiv
Top 0.5%
1.3%
Show abstract

Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.

11
AI Video Analysis of Psychomotor Performance in EMS Education: Agreement With Human Evaluators Across Three Skills

Otte, J. H.; Cartagena, A.

2026-08-31 medical education 10.64898/2026.08.26.26361437 medRxiv
Top 0.5%
1.1%
Show abstract

Background. A primary constraint on the capacity of EMS programs to meet industry demand is psychomotor instruction and verification, requiring direct observation of each student by a qualified evaluator. Whether AI video analysis can relieve it is untested; none has been applied to EMS skill examination or compared with human examiners. Objective. To quantify human EMS evaluator inter-rater reliability and evaluate an AI video-analysis platform against it. Methods. In a prospective, fully crossed study, five certified EMS evaluators and an AI platform independently scored identical video-recorded EMT performances of cervical collar application (n=15), bag-valve-mask (BVM) ventilation (n=14), and medical assessment (n=15) on dichotomous checklists with critical-failure criteria. Agreement was assessed at item, score, and decision levels using Fleiss' kappa, Krippendorff's alpha, Gwet's AC1, and ICC(2,1)/ICC(2,k). Results. Human item agreement was moderate (kappa 0.409 to 0.467), as was single-rater reliability (ICC(2,1) 0.539 to 0.694), against good panel reliability (ICC(2,k) 0.854 to 0.919). Recorded pass/fail agreement was fair (kappa 0.297 to 0.388) and critical-failure agreement near zero for two skills (kappa 0.028, 0.119). AI alignment tracked rubric observability rather than task complexity: r = 0.857 (collar, exceeding every human), -0.173 (BVM), 0.664 (medical), and it was most lenient on two skills. Conclusions. Human evaluators are an imperfect standard, especially on critical failures. The AI was a legitimate additional rater where checklist items were discrete and visually verifiable, but not where credit required judging continuous quantities such as ventilation rate, volume, or suction duration. Defensible uses are formative and archival, not summative. These results reflect an early, non-specialist configuration: a baseline, not a limit.

12
Feasibility study of gait analysis using a new Wearable Force Plate

Sanz Morere, C. B.; Garrido-Lopez, G.; Hayase, M.; Rueda, J.; An, Q.; Shimoda, S.; Moreno, J. C.; Navarro, E.

2026-09-02 rehabilitation medicine and physical therapy 10.64898/2026.08.30.26361786 medRxiv
Top 0.8%
0.8%
Show abstract

Static force plates (FP) are the gold standard for measuring ground reaction forces (GRF) and computing joint moments through inverse dynamics in gait analysis. However, they are restricted to controlled environments, and the number of steps analyzed is limited by the plates embedded in the floor. To address these limitations, portable solutions such as sensorized insoles, socks, or shoes have emerged. Yet, creating wearable systems capable of measuring three-dimensional GRF in real-world conditions remains challenging. Current sensorized shoes often incorporate thick sensors (up to 2 cm), reducing usability and limiting their application in pathological populations or dynamic tasks like running. This study evaluates the usability of ShokacShoes, a novel sensorized shoe integrating three thin, three-dimensional force sensors, and explores its potential as a Wearable Force Plate (WFP). Eight healthy participants performed slow, natural, and fast walking using two insole configurations. Force and temporal metrics were derived from WFP and FP data. Results indicate that WFP enables accurate step segmentation and detects significant effects of speed and insole type on temporal and force metrics, confirming its reliability under different walking conditions. Comparisons with FP revealed differences in force metrics and signal morphology, though temporal parameters remained consistent. These results are likely due to sensor quantity and positioning. Thereby, ShokacShoes represent a valid solution capable of measuring three-dimensional forces within commercial footwear. Future work will focus on validating the applicability of a new version of ShokacShoes against gold-standard FP in a comprehensive validation study involving diverse real-world scenarios and pathological conditions.

13
The use of computerised testing to assess cognitive performance in people with HIV in South Africa

Edmond, E. C.; Dreyer, A. J.; Winston, A.; Khoo, S. H.; Joska, J.; Nightingale, S.

2026-08-31 hiv aids 10.64898/2026.08.27.26361083 medRxiv
Top 1%
0.5%
Show abstract

Background Computerised cognitive testing may address the global challenge in identifying cognitive changes in people living with HIV scalably and affordably. We assessed a computerised battery (CB) of cognitive tests, in a prospective cohort (CONNECT) of people with HIV in a low-income peri-urban area of Cape Town, South Africa during a national programmatic switch from efavirenz- to dolutegravir-based antiretroviral therapy (ART). Methods We recruited 170 people with HIV and 91 people without HIV (controls) (140[82%] and 41[45%] followed up). The CB and gold-standard pen&paper cognitive testing (P&P) were performed at both timepoints. Technology familiarity/use questionnaire data were also collected. We compared performance in detecting lower group-level cognitive performance associated with efavirenz treatment. Furthermore, the CB was compared to P&P in classifying individuals with low cognitive performance, correlation of global test scores and domain-level scores between batteries, and practice effects between timepoints. Exploratory principal component analysis was also performed. Results People with HIV on efavirenz at baseline had lower performance on the computerised battery than controls, {Delta}T=2.6, p=0.0047. This difference was lost after switching to dolutegravir-based ART at follow-up. CB and P&P global T were moderately correlated (R2=0.203, p<0.001), and the CB performed moderately in classification of low cognitive performance against the gold standard (AUC 0.70, sensitivity 0.52, specificity 0.76, PPV 0.40, and NPV 0.84). Selecting the first three principal components improved both classification of low cognitive performance (AUC 0.77) and correlation strength with P&P global T (R2=0.3, p<0.001). The CB did not show practice effects. Most participants owned a mobile phone (95%, 85.9% of these smartphones). Performance was better in smartphone owners ({Delta}T=1.8) and computer owners (23%, {Delta}T=1.8). Conclusions Delivering computerised cognitive testing was feasible in this low-income southern African setting. The CB showed reasonable construct validity (detecting known lower cognitive performance associated with efavirenz-ART) and may detect broad cognitive characteristics such as processing speed and accuracy. However, correlation of CB results with gold standard P&P testing was low-moderate and may limit its applicability as a diagnostic tool. This might be improved by including a wider range of cognitive domains tested in the CB, or data driven analysis. Brief CBs may fulfil an initial screening role, followed by more detailed clinical assessment.

14
CHARMS and PROBAST+AI: an updated template for Data Extraction and Risk of Bias Assessment in systematic reviews of prediction models

Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.

2026-08-31 cardiovascular medicine 10.64898/2026.08.26.26361189 medRxiv
Top 1%
0.5%
Show abstract

Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.

15
Usability, acceptability and feasibility of continuous glucose monitoring among children and adolescents with type 1 diabetes in Kenya

Amolo, P.; Mungai, L.; Karume, A. K.; Kibugi, J.; Mwende, W.; Botella, N.; Haldane, C.; Kamau, Y.; Marban-Castro, E.

2026-09-01 endocrinology 10.64898/2026.08.27.26361447 medRxiv
Top 1%
0.4%
Show abstract

Introduction Continuous Glucose Monitoring (CGM) is considered standard care in high-income countries. There is, however, limited published evidence on CGM use in low- and middle-income countries. The purpose of this study was to assess the usability, acceptability, and feasibility of CGM use among people living with type 1 diabetes (T1D) and caregivers in a low-resource setting. Research Design and Methods This prospective study conducted at the Kenyatta National Hospital purposively enrolled persons aged 4-25 years who had been on management for T1D for at least six months, and caregivers of those under 18 years. Fourty youth living with T1D used CGM for three months in place of self monitoring of blood glucose (SMBG). The System Usability Scale (SUS), a Theoretical Framework of Acceptability-based questionnaire, the Diabetes Distress Scale (DDS), the Glucose Monitoring Satisfaction Survey (GMSS), and a feasibility survey were administered. Outcomes were summarized descriptively, including means, medians, and frequencies using R statistical software. Results The median SUS score was 98.8 (IQR 92.5-100.0). Acceptability was high, and the median total GMSS score improved from 3.73 to 4.73. Among adolescents and adults, the median overall DDS score reduced from 1.54 to 1.36, with reductions in scores in all domains, except for hypoglycemia distress which increased, and physician distress which remained low. Among caregivers, the median overall DDS score declined from 2.05 (moderate distress) to 1.90 (low distress), with modest reductions in teen management and parent-teen relationship distress and a slight increase in personal distress. Median CGM active wear time was 89%. Conclusion This study comprehensively evaluated CGM across usability, acceptability, and feasibility outcomes, with the findings supporting the integration of CGM into routine diabetes management in low-resource settings. The short follow-up period, however, may not capture changing perceptions or long-term adherence.

16
Glaucoma and Diabetes Mellitus: A Comparative Evaluation of Comorbid Effect on Tear Quantity among Patients in Owerri, Imo State, Nigeria.

Chukwuoha, C. M.; Ovenseri-Ogbomo, G.; Azuamah, Y. C.; Odimegwu, N. E.; Obioma-Elemba, J. E.; Ugwoke, G.; Nkeremuzor, E. C.; Eronini, Y.; Ikoro, N. C.; Esenwah, E. C.

2026-09-02 ophthalmology 10.64898/2026.08.30.26361782 medRxiv
Top 1%
0.4%
Show abstract

Abstract Objective: Glaucoma is a chronic disorder that impairs ocular health and may exacerbate ocular surface disease leading to tear film instability, dry eye symptoms and decreased quality of life. This study compared changes in tear quantity among glaucoma subjects living with and without diabetes mellitus, attending an eye clinic in Nigeria. Methods: A comparative cross sectional research design was used. 157 subjects which comprised 74 glaucoma subjects living with diabetes mellitus and 83 glaucoma subjects living without diabetes mellitus participated in the study. Tear quantity assessment included the Schirmer I test and tear meniscus height (TMH) measurement. Descriptive statistics, independent samples t-test and Chi-square test were used to examine the data at 0.05 level of significance. Results: Glaucoma subjects living with diabetes mellitus showed substantially decreased tear production (11.4 +/- 6.8 mm) compared with glaucoma subjects living without diabetes mellitus (19.6 +/- 9.6 mm; p < 0.001). Tear meniscus height in glaucoma subjects living with diabetes mellitus (0.8 +/- 0.3 mm) was significantly greater than in subjects living without diabetes mellitus (0.7 +/- 0.3 mm; p = 0.034). Conclusion: Diabetes mellitus dramatically deteriorates the ocular surface function in glaucoma subjects by decreasing tear production, altering the tear meniscus height and increasing the severity of ocular surface symptoms. Routine glaucoma care, especially in patients with diabetes mellitus, should include a full ocular surface evaluation including Schirmer I test, TBUT, TMH, and OSDI assessment to allow early detection and management of ocular surface disease, better treatment adherence, and improved visual outcomes. Keywords: Glaucoma, Diabetes Mellitus, Tear production, Tear Meniscus Height, Ocular Surface Disease.

17
Empowering adults to manage their hearing loss: assessing the benefits of user-controlled, smartphone-connected hearing aids.

Maidment, D. W.; Habib, A.; Gomez, R.; Benton, C.; Ferguson, M. A.

2026-09-03 otolaryngology 10.64898/2026.08.30.26361775 medRxiv
Top 1%
0.4%
Show abstract

The availability of hearing aids that can connect wirelessly to smartphone technologies via Bluetooth has grown exponentially in recent years. However, there is limited evidence assessing the benefits of user-adjustability afforded by these devices. This study aimed to assess the benefits of smartphone-connected hearing aids and an accompanying application (or app) in new and existing hearing aid users. In this single-centre, prospective, observational study, 44 adult hearing aid users (14 new and 30 existing) were recruited. Participants were fitted bilaterally with smartphone-connected hearing aids that could be adjusted by the user via an app. Self-reported outcome measures were collected at fitting and after seven-weeks of using the device in everyday life. For both new and existing hearing aid users, significant improvements in social participation, hearing-related fatigue, and hearing aid benefit and satisfaction were found. For existing hearing aid users, all outcomes were significantly better for the smartphone-connected hearing aids plus app in comparison to their existing hearing aids that did not connect to a smartphone, all with moderate-to-large clinical effect sizes (d> .6). User-controllability via the app was identified as the key benefit, and most participants (68%) reported that the app met their needs 'extremely' or 'very well'. These results suggest that, when used in conjunction with an app, smartphone-connected hearing aids can improve hearing outcomes due to greater user-controllability to improve listening. Thus, smartphone-connected hearing aids have the potential to facilitate patient-centred care, empowering the individual to successfully manage their hearing loss.

18
Evaluating Clinical Foundation Models for Early Alzheimer's Disease and Related Dementia Prediction from Longitudinal EHRs

Farzana, S.; Arian, A.; Rundek, T.; Desvarieux, M.; Ahsan, H.

2026-09-03 health informatics 10.64898/2026.09.01.26361933 medRxiv
Top 1%
0.4%
Show abstract

Early identification of Alzheimer's disease and related dementias (ADRD) remains challenging despite its importance for timely intervention, management of modifiable risk factors, and care planning. We developed and evaluated ADRD onset prediction models using longitudinal electronic health records (EHRs) from the All of Us Research Program at clinically meaningful lead times of 6, 12, 24, and 36 months before diagnosis, benchmarking interpretable count-based representations against four publicly available pretrained clinical foundation models (CLMBR-T, GPT-style, LLaMA-style, and Mamba) across multiple ADRD phenotype definitions. Count-based models consistently achieved the highest discrimination and calibration across all cohorts and prediction horizons. Predictive performance declined with increasing lead time for all approaches; however, the performance gap between count-based and pretrained representations progressively narrowed, with foundation models achieving comparable AUROC of 0.719 (compared to the AUROC of 0.738 of count-based model) at the 36-month horizon while providing higher sensitivity and F1 scores under a fixed operating threshold. External validation with zero-shot evaluation on UChicago EHRs exhibited limited generalizability for count-based and pretrained clinical foundation model based representations. These findings demonstrate that transparent count-based EHR representations remain the strongest overall approach for ADRD onset prediction, while pretrained clinical foundation models provide complementary advantages for long-term risk identification and establish a benchmark for evaluating transferable clinical representations in temporal ADRD risk prediction.

19
Longitudinal Tracking and Construct Validity of a Single-Item Physical Activity Measure in the Womens Healthy Ageing Project

Corcoran, D.; Szoeke, C.; Apostolopoulos, V.; Feehan, J.

2026-08-31 public and global health 10.64898/2026.08.27.26361567 medRxiv
Top 1%
0.3%
Show abstract

This study aimed to quantify the longitudinal tracking and cross-sectional construct validity of a single-item questionnaire measuring recreational physical activity frequency (RPAF) in the Womens Healthy Ageing Project. At baseline, 474 participants aged 45-55 reported RPAF from 1993 to 2014. Longitudinal tracking of the RPAF item was assessed as a consecutive-wave and baseline-referenced measure using linear weighted kappa (LWK), Spearman correlations, exact agreement and within-one-category agreement. Construct validity in the form of convergent and known-group validity was assessed using the International Physical Activity Questionnaire (IPAQ) leisure activity domains, Short Form 36 physical function (SF-36-PF) subscale, Timed Up and Go (TUG), hand grip strength (HGS) and waist-to-height ratio (WHtR). 474 participants provided baseline RPAF data. Pairwise longitudinal samples ranged from 176 to 459 across the study. Consecutive-wave LWK ranged from 0.38 to 0.49, and Spearman correlations ranged from 0.44 to 0.57. Exact and within-category agreement ranged from 41.4%-50.8% and 72.0%-79.0%. Baseline-referenced LWK ranged from 0.22 to 0.47, with Spearman correlations of 0.29 to 0.56. RPAF correlated with total IPAQ leisure score (rs = 0.60), IPAQ walking score (rs = 0.58), SF-36-PF (rs = 0.33) and TUG score (rs = -0.25). No significant correlation was identified between RPAF, HGS or WhTR. RPAF discriminated known groups for WHO guideline-sufficient activity, SF-36-PF, and TUG fall risk. The RPAF item demonstrated fair-to-moderate agreement in consecutive waves, with weaker baseline-referenced tracking. Cross-sectional validity was highest with total IPAQ leisure activity. The item may provide a pragmatic measure for RPAF in womens cohort studies.

20
Acceptability, feasibility, quality of life and diabetes distress score outcomes: A pragmatic randomised clinical trial on continuous glucose monitoring for people with type 1 diabetes

Marban-Castro, E.; Muhwava, L.; Girdwood, S.; Kemp, T.; Freitas, J.; Kamau, Y.; Otieno, M.; Akach, D.; Morato, A.; Sanz, S.; Fiechter, V.; Erkosar, B.; Watson, M.; Vetter, B.; Haldane, C.; Shilton, S.; Rheeder, P.; Dave, J. A.; Carrihill, M.; Karsas, M.

2026-08-31 endocrinology 10.64898/2026.08.26.26361479 medRxiv
Top 1%
0.3%
Show abstract

Introduction: Continuous glucose monitoring (CGM) offers an advancement over traditional self-monitoring of blood glucose (SMBG) for people living with type 1 diabetes (T1D). However, evidence on the acceptability and feasibility of different CGM use cases in African populations remains limited. Methods: This was a pragmatic three-arm, randomised controlled trial on CGM conducted among people living with T1D in three public healthcare clinics in South Africa. Participants were assigned to Arm 1 (continuous CGM), Arm 2 (periodic CGM), or Arm 3 (SMBG). Diabetes education was provided at all study visits. Feasibility was assessed by adherence to CGM use and through the Glucose Monitoring Satisfaction Survey (GMSS). Diabetes distress was measured by the Diabetes Distress Scale (DDS), health-related quality of life (HRQoL) by the EQ-5D scales, and acceptability using the Theoretical Framework of Acceptability (TFA). Surveys were collected on paper and transferred to OpenClinica. Analyses were performed in R. The trial was registered in the Clinical Trials Registry (NCT05944718) on July 13, 2023. Results: A total of 83 participants were included in Arm 1, 85 in Arm 2, and 80 in Arm 3. CGM mean active time was 55% in Arm 1 versus 69% in Arm 2. The proportion of participants meeting the [&ge;]70% active time threshold was higher in Arm 2 (52%) than in Arm 1 (34%). Diabetes' distress declined across arms during the intervention period, with no significant difference between arms; distress increased slightly six months post-intervention but remained below baseline. At 6 months, glucose monitoring satisfaction was significantly higher in both CGM arms than in the SMBG arm, and satisfaction increased over time in CGM arms. Health-related quality of life remained stable across arms during the intervention period with no significant difference between arms. High acceptability was observed in both CGM arms, with higher ratings in the periodic arm. Conclusions: CGM was acceptable to people living with type 1 diabetes and feasible to use in public-sector clinics in South Africa, with high acceptability under continuous and periodic use. Health-related quality of life remained stable across arms, and diabetes-related distress declined, during the intervention period, across arms. Glucose monitoring satisfaction rose significantly in both CGM arms compared to SMBG. Periodic CGM might be a promising and potentially more scalable option than continuous use for public-sector care.